Global Trajectory-Aware Multi-camera Multi-target Tracking
FAN Yuteng1,2, WANG Qiang3, ZHEN Yihao1,2, CHEN Xiai1, MA Chi4, FAN Huijie1
1. State Key Laboratory of Robotics and Intelligent Systems, Shen-yang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016; 2. School of Computer Science and Technology, University of Chinese Academy of Sciences, Beijing 100049; 3. Liaoning Key Laboratory of Integrated Automation for Equipment Manufacturing, Shenyang University, Shenyang 110096; 4. School of Computer Science and Engineering, Huizhou University, Huizhou 516007
Abstract:In existing global tracking methods, global trajectories are mainly treated as collections of historical target features. However, trajectory-level identity semantics are not explicitly modeled, and identity consistency constraints are not imposed. Therefore, the long-term identity consistency of targets across time and space cannot be fully exploited. To address these issues, a global trajectory-aware multi-camera multi-target tracking method termed GTAT is proposed. First, for each global trajectory, a trajectory-level representation with long-term identity semantics is aggregated by the global trajectory semantic encoding module. Then, global trajectory-level identity semantics are injected into target association features within the temporal window by the global trajectory-aware enhancement module through a cross-attention mechanism. Thus, trajectory-level contextual information can be explicitly perceived by target-level representations. Finally, a global trajectory consistency loss is designed to constrain the matching relationship between trajectory-level representations and target-level association features. Consequently, the ability of global trajectory representations to model identity consistency is improved. The identity discriminability of association features is also enhanced. Experiments demonstrate that GTAT achieves superior performance on six public multi-camera multi-target tracking datasets.
[1] Li Y J, Weng X S, Xu Y, et al. Visio-temporal attention for multi-camera multi-target association[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Washington, USA: IEEE, 2021: 9814-9824. [2] Li D G, Wei X, Hong X P, et al. Infrared-visible cross-modal person re-identification with an X modality[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(4): 4610-4617. [3] Yogamani S, Hughes C, Horgan J, et al. WoodScape: a multi-task, multi-camera fisheye dataset for autonomous driving[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Washington, USA: IEEE, 2019: 9307-9317. [4] Ma Z H, Wei X, Hong X P, et al. Bayesian loss for crowd count esti-mation with point supervision[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Washington, USA: IEEE, 2019: 6141-6150. [5] 濮志远,罗素云.复杂交通场景下的目标检测方法[J].信息与控制, 2025, 54(4): 632-643. (Pu Z Y, Luo S Y.Object detection method in complex traffic scenarios[J]. Information and Control, 2025, 54(4): 632-643.) [6] 陈涛,肖杰,张雨飞,等.多相似度矩阵融合的多目标跟踪算法[J].小型微型计算机系统, 2026, 47(2): 428-434. (Chen T, Xiao J, Zhang Y F, et al. Multi-object tracking algorithm based on multi-similarity matrix fusion[J]. Journal of Chinese Computer Systems, 2026, 47(2): 428-434.) [7] Xu Y L, Liu X B, Qin L, et al. Cross-view people tracking by scene-centered spatio-temporal parsing[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2017, 31(1): 4299-4305. [8] Han R Z, Feng W, Zhao J W, et al. Complementary-view multiple human tracking[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(7): 10917-10924. [9] 赵永祥,张国庆,李灯华,等.一种用于监测奶牛的高精度跟踪定位方法[J].信息与控制, 2025, 54(1): 137-149. (Zhao Y X, Zhang G Q, Li D H, et al. A high-precision tracking and localization method for monitoring cows[J]. Information and Control, 2025, 54(1): 137-149) [10] 乔羽,范慧杰,付生鹏,等.高空无人机跨场景多目标跟踪方法[J].小型微型计算机系统, 2025, 46(8): 1993-1999. (Qiao Y, Fan H J, Fu S P, et al. High-altitude unmanned aerial vehicle multi-target tracking method[J]. Journal of Chinese Computer Systems, 2025, 46(8): 1993-1999.) [11] Zhang Y F, Sun P Z, Jiang Y, et al. ByteTrack: multi-object tracking by associating every detection box[C]//Proceedings of the 17th European Conference on Computer Vision. Berlin, Germany: Springer, 2022: 1-21. [12] Specker A.OCMCTrack: online multi-target multi-camera tracking with corrective matching cascade[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Washington, USA: IEEE, 2024: 7236-7244. [13] Quach K G, Nguyen P, Le H, et al. DyGLIP: a dynamic graph model with link prediction for accurate multi-camera multiple object tracking[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2021: 13779-13788. [14] He Y H, Wei X, Hong X P, et al. Multi-target multi-camera tracking by tracklet-to-target assignment[J]. IEEE Transactions on Image Processing, 2020, 29: 5191-5205. [15] Gan Y Y, Han R Z, Yin L Q, et al. Self-supervised multi-view multi-human association and tracking[C]//Proceedings of the 29th ACM International Conference on Multimedia. New York, USA: ACM, 2021: 282-290. [16] Feng W, Wang F F, Han R Z, et al. Unveiling the power of self-supervision for multi-view multi-human association and tracking[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(1): 351-368. [17] Liu Z H, Shang Y Y, Li T M, et al. Robust multi-drone multi-target tracking to resolve target occlusion: a benchmark[J]. IEEE Transactions on Multimedia, 2023, 25: 1462-1476. [18] Hao S Y, Liu P Y, Zhan Y B, et al. DIVOTrack: a novel dataset and baseline method for cross-view multi-object tracking in diverse open scenes[J]. International Journal of Computer Vision, 2024, 132(4): 1075-1090. [19] Amosa T I, Sebastian P, Izhar L I,et al. Multi-camera multi-object tracking: a review of current trends and future advances[J/OL]. Neurocomputing, 2023, 552. https://doi.org/10.1016/j.neucom.2023.126558. [20] Zhen Y H, Xu M Y, Wang Q, et al. GMT: effective global framework for multi-camera multi-target tracking[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2026: 28201-28210. [21] Fan H J, Qiao Y, Zhen Y H, et al. All-day multi-camera multi-target tracking[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2025: 16892-16901. [22] 范慧杰,郁航,赵颖畅,等.可见光红外跨模态行人重识别方法综述[J].信息与控制, 2025, 54(1): 50-65. (Fan H J, Yu H, Zhao Y C, et al. Review of visual-infrared cross-modal person re-identification methods[J]. Information and Control, 2025, 54(1): 50-65.) [23] Yu F, Wang D Q, Shelhamer E, et al. Deep layer aggregation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2018: 2403-2412. [24] Duan K W, Bai S, Xie L X, et al. CenterNet: keypoint triplets for object detection[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Washington, USA: IEEE, 2019: 6568-6577. [25] Shao S, Zhao Z J, Li B X, et al. CrowdHuman: a benchmark for detecting human in a crowd[EB/OL].[2026-05-12]. https://arxiv.org/abs/1805.00123. [26] Fleuret F, Berclaz J, Lengagne R, et al. Multicamera people tra-cking with a probabilistic occupancy map[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2008, 30(2): 267-282. [27] Xu Y L, Liu X B, Liu Y, et al. Multi-view people tracking via hierar-chical trajectory composition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2016: 4256-4265. [28] Chavdarova T, Baqué P, Bouquet S, et al. WILDTRACK: a multi-camera HD dataset for dense unscripted pedestrian detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2018: 5030-5039. [29] 范晓雨,刘明华,李文静.无人机场景下多层次特征变换的池化Transformer跟踪算法[J].小型微型计算机系统, 2025, 46(11): 2578-2585. (Fan X Y, Liu M H, Li W J.Multi-level feature transformation of pooling transformer for UAV object tracking[J]. Journal of Chinese Computer Systems, 2025, 46(11): 2578-2585.) [30] Bernardin K, Stiefelhagen R. Evaluating multiple object tracking performance: the CLEAR MOT metrics[J/OL]. EURASIP Journal on Image and Video Processing, 2008, 2008. https://link.springer.com/content/pdf/10.1155/2008/246309.pdf [31] Luiten J, Ošep A, Dendorfer P, et al. HOTA: a higher order me-tric for evaluating multi-object tracking[J]. International Journal of Computer Vision, 2021, 129(2): 548-578. [32] Ristani E, Solera F, Zou R, et al. Performance measures and a data set for multi-target, multi-camera tracking[C]//Proceedings of the European Conference on Computer Vision Workshops. Berlin, Germany: Springer, 2016: 17-35. [33] Ye M, Shen J B, Lin G J, et al. Deep learning for person re-identification: a survey and outlook[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(6): 2872-2893. [34] Wieczorek M, Rychalska B, Dabrowski J.On the unreasonable effec-tiveness of centroids in image retrieval[C]//Proceedings of the 28th International Conference on Neural Information Processing. Berlin, Germany: Springer, 2021: 212-223.